Papers with automatic processing
Toward Implicit Reference in Dialog: A Survey of Methods and Data (2022.aacl-main)
Copied to clipboard
| Challenge: | In natural language, speakers often leave out information that is understood by the other party through the shared context. |
| Approach: | They propose to use omitted entities as implicit references in dialogs to improve language processing. |
| Outcome: | The proposed method is based on a set of experiments which show that the proposed method has a high level of accuracy and is a success. |
At the Crossroad of Cuneiform and NLP: Challenges for Fine-grained Part-of-speech Tagging (2024.lrec-main)
Copied to clipboard
| Challenge: | cuneiform texts are dominated by multiple languages and language families . the most dominant language written in cuniform is the Semitic Akkadian . existing cnl models are not suitable for digital editions of Akkadi . |
| Approach: | They focus on letters written in the Semitic Akkadian, a cuneiform language dominated by cuniform texts . they propose to use pre-trained embeddings, sentence segmentation and cnl to fine-tune language models . |
| Outcome: | The dominant language written in cuneiform is the Semitic Akkadian . the paper examines the input material and tries to initiate a discussion about best-practices . |
Processing Language Resources of Under-Resourced and Endangered Languages for the Generation of Augmentative Alternative Communication Boards (2020.lrec-1)
Copied to clipboard
| Challenge: | Under-resourced and endangered or small languages yield problems for automatic processing and exploiting because of the small amount of available data. |
| Approach: | They propose an approach using enriched linguistic research data to create communication boards commonly used in alternative augmentative communication (AAC) using lexical analysis and rich annotation, the boards can be imported into various AAC software. |
| Outcome: | The proposed approach uses lexical analysis and rich annotations to create communication boards commonly used in alternative augmentative communication (AAC) The created boards can be imported into various AAC software and are available under the CC BY-NC-SA 4.0 (public) license. |
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification. (2022.lrec-1)
Copied to clipboard
Rémi Uro, David Doukhan, Albert Rilliard, Laetitia Larcher, Anissa-Claire Adgharouamane, Marie Tahon, Antoine Laurent
| Challenge: | Existing methods for creating diachronic corpus of voices are based on speaker characteristics and require human intervention. |
| Approach: | They propose to use a semi-automatic pipeline to create a diachronic corpus of voices balanced for speaker’s age, gender and recording period, according to 32 categories. |
| Outcome: | The proposed method cut down on manual annotations by ten and provides high quality speech for most of the selected excerpts. |
Preserving Semantic Information from Old Dictionaries: Linking Senses of the ‘Altfranzösisches Wörterbuch’ to WordNet (2020.lrec-1)
Copied to clipboard
| Challenge: | Historical dictionaries of the pre-digital period are important resources for the study of older languages. |
| Approach: | They propose to use printed dictionaries to create a more easily accessible and more sustainable lexical database by automating the conversion process. |
| Outcome: | The ‘Altfranzösisches Wörterbuch’, an Old French dictionary published from 1925 onwards, shows how the printed dictionaries can be turned into a more easily accessible and more sustainable lexical database. |
The Hebrew Essay Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Annotated corpus of argumentative essays authored by prospective higher-education students . corpus includes essays by native speakers and essays by non-native speakers . |
| Approach: | They propose to use an annotated corpus of Hebrew argumentative essays to analyze non-native language use. |
| Outcome: | The proposed corpus includes essays by native speakers and essays authored by non-native speakers with three different native languages. |
Marking Irony Activators in a Universal Dependencies Treebank: The Case of an Italian Twitter Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing annotations for irony are difficult, and the recognition of it is difficult due to its polarity. |
| Approach: | They propose a fine-grained annotation scheme centered on irony that highlights the tokens responsible for its activation and their morpho-syntactic features. |
| Outcome: | The proposed scheme highlights the tokens responsible for irony activation and their morpho-syntactic features. |
Parallel Corpora in Mboshi (Bantu C25, Congo-Brazzaville) (L18-1)
Copied to clipboard
Annie Rialland, Martine Adda-Decker, Guy-Noël Kouarata, Gilles Adda, Laurent Besacier, Lori Lamel, Elodie Gauthier, Pierre Godard, Jamison Cooper-Leavitt
| Challenge: | BULB project aims to provide tools to language documentation and description for unwritten languages . language-based technologies are needed to support the collection of data and to provide linguistic documentation for the languages. |
| Approach: | This paper presents multimodal and parallel data collections in Mboshi, as part of the French-German BULB project. |
| Outcome: | The proposed data collection includes pictures and videos documenting social practices, agriculture, wildlife and plants. |